Phase 3: Machine Learning Lesson 1 of 6

How Machines Learn:
The Core Idea

This is the lesson where everything clicks. You have seen what data is and how it is structured. Now you will see what a machine actually does with that data to get smarter over time.

You will learn
The machine learning loop in plain English
What "training" a model actually means
How gradient descent finds the best answer
The difference between parameters and hyperparameters

What does it mean to learn?

When you were learning to ride a bike, nobody wrote you a manual with instructions like "apply 2.3 Newtons of force to the left pedal at exactly 38 degrees." That would have been useless. Instead, you got on the bike, tried, fell, adjusted, tried again, and slowly your body figured out the balance. Trial, error, adjustment, repeat.

Machine learning works the same way. You do not program rules into the model. You give it data, let it make predictions, show it how wrong those predictions were, and let it adjust. Do that thousands or millions of times and the model gets good at the task. The machine is learning from feedback, exactly like you learned to ride that bike.

The difference is scale. A human might need a few hundred attempts to learn a skill. A machine can run through millions of examples in minutes.

"Machine learning is a field of study that gives computers the ability to learn without being explicitly programmed."

Arthur Samuel, 1959 — the person who coined the term "machine learning"

The learning loop

Every machine learning model, from the simplest linear regression to the largest language model, follows the same basic cycle. Understanding this loop is the most important conceptual foundation in this entire phase.

📦
Data in
Feed a batch of training examples into the model
🔮
Predict
The model makes a prediction for each example
📏
Measure error
Compare prediction to the real answer, calculate the loss
↩️
Adjust
Tweak the model's internal numbers to reduce the error
🔁
Repeat
Run this cycle thousands of times until the error is small

That cycle runs over and over, adjusting the model slightly each time, until the predictions are good enough. The number of times you run the full loop is called an epoch. Running for 10 epochs means the model has seen the full training dataset 10 times.

What the model is actually adjusting

Inside every machine learning model there are numbers called parameters (also called weights and biases). These are the knobs the model turns during training. At the start, they are set randomly. By the end of training, they hold the learned pattern.

Think of a linear model predicting house prices. Its prediction might look like this:

ConceptLinear model
price = (w1 × size) + (w2 × bedrooms) + (w3 × location) + b

# w1, w2, w3 are the weights (parameters the model learns)
# b is the bias (another learned parameter)
# At the start: w1=0.1, w2=0.3, w3=-0.2, b=50
# After training: w1=180, w2=25000, w3=42000, b=-15000

The model starts with random values for all those weights. It is terrible at first. After training, those weights encode the actual relationship between features and price that the model discovered in the data. Training is nothing more than finding the right values for these numbers.

How does it know which way to adjust?

This is where gradient descent comes in. Gradient descent is the algorithm that figures out how to adjust the weights to reduce the error. The intuition is simple, even if the full mathematics runs deeper.

Gradient descent: rolling toward the lowest point
model weights error (loss) start minimum

Imagine the curve as a valley. The ball starts somewhere random on the slope. Gradient descent figures out which way is downhill and takes a small step in that direction. Repeat enough times and the ball settles at the lowest point, where the error is smallest. That lowest point is your trained model.

The size of each step is called the learning rate. Too large and you overshoot the bottom and bounce around. Too small and training takes forever. Choosing a good learning rate is one of the most important practical decisions in machine learning.

Parameters vs hyperparameters

There are two kinds of numbers in a machine learning model and it is worth being clear about the difference because the terms come up constantly.

Parameters
w1, w2, b
The numbers the model learns during training. They start random and get adjusted through gradient descent. You never set these yourself. The model finds them.
Hyperparameters
learning_rate, epochs, n_trees
The settings you choose before training starts. How fast should the model learn? How many trees in the forest? These are not learned; they are decisions you make.
A useful way to remember it

Parameters are what the model learns. Hyperparameters are what you decide before learning begins. A chef does not choose how the dish tastes while eating it. They choose the recipe, the oven temperature and the cooking time before the dish goes in. Those choices are the hyperparameters. What comes out is shaped by the parameters the process discovers.

Your first model in Scikit-learn

All of this theory becomes much more concrete once you see the loop in code. Scikit-learn makes this remarkably clean. Here is a full machine learning pipeline in about 10 lines.

PythonThe full ML loop in Scikit-learn
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error
import numpy as np

# 1. Prepare your features and label
X = df[['size', 'bedrooms', 'location_score']]
y = df['price']

# 2. Split into training and test sets
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42)

# 3. Create the model (a hyperparameter choice: which model?)
model = LinearRegression()

# 4. Train: the loop runs inside here
model.fit(X_train, y_train)

# 5. Predict on unseen data
predictions = model.predict(X_test)

# 6. Measure the error
error = mean_squared_error(y_test, predictions)
print(f"Mean squared error: {error:.2f}")

Notice that you never write the learning loop yourself. model.fit() runs the entire gradient descent process for you. Scikit-learn handles all the maths internally. Your job is to prepare clean data, choose a sensible model, and then evaluate whether the result is good enough.

Think of it this way

Calling model.fit() is like hiring an experienced chef and handing them a recipe book full of examples. The chef studies all the examples, figures out the patterns, and develops their own internal skills. You do not watch every step of that process. When they are done, you call model.predict() and they cook you a meal based on everything they learned.

What training cannot do

It is worth being honest about the limits of this process. The learning loop optimises the model to perform well on its training data. That is not the same as understanding the world. A model that has learned from historical house prices knows nothing about what a house physically is. It knows which numbers tend to go together.

This is why a model trained on data from one city can perform poorly in another. It has learned the patterns in what it saw, not the underlying rules of reality. Keeping this distinction in mind will save you from over-trusting your models and will help you diagnose problems when they arise.

Lesson Activity · No code required
Draw the Loop
Before you write a single line of ML code, make sure this loop is automatic in your head. This exercise takes 5 minutes and will pay back every time you need to debug a model.
01 Close this lesson and draw the 5-step machine learning loop from memory. Use your own words for each step, not the labels from this page.
02 Pick a real-world task: predicting whether a patient has diabetes. Write one sentence for what happens at each of the 5 steps for that specific problem.
03 List two parameters (things the model learns) and two hyperparameters (things you decide) for a model predicting diabetes risk. Bring your answers to the live session.
Your Notes
Studying independently? Write your thoughts or answers below. Notes save automatically to your browser.

Reflect

Before you move on

No right answers here. These questions are for you.

When we say a model "learns," what is actually changing inside it during training?

The model's parameters (weights and biases) are being adjusted. Starting from random values, each training example produces a prediction. The error between that prediction and the correct answer is calculated, and the parameters are nudged slightly in the direction that reduces the error. After millions of such adjustments, the parameters settle into values that work well across the data.

Name a task where writing explicit rules would be easy, and one where machine learning is clearly the better approach. What makes the difference?

Calculating a tax bill follows clear rules that can be written out precisely; a rule-based system works fine. Recognising a face in a photo does not: the number of edge cases, lighting conditions, and angles is too vast to enumerate. The key difference is whether the problem has a definable, finite rule set or whether the pattern is too complex and variable to specify by hand.

Gradient descent finds a minimum by always stepping downhill. What is the risk with this approach, and why does it still work so well in practice?

The risk is getting stuck in a local minimum, a valley that is not the lowest point globally. In simple examples this is a real problem. In practice, neural networks operate in millions of dimensions, and research suggests that most local minima in high-dimensional spaces are actually close in quality to the global minimum. Techniques like random restarts, momentum, and adaptive learning rates also help escape shallow traps.

Progress
Done with this lesson?
Mark it complete to track your progress.